Back

Nature Methods

Springer Science and Business Media LLC

All preprints, ranked by how well they match Nature Methods's content profile, based on 385 papers previously published here. The average preprint has a 0.38% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
ShareLoc - an open platform for sharing localization microscopy data

Bai, J.; Ouyang, W.; Singh, M. K.; Leterrier, C.; Barthelemy, P.; Barnett, S. F. H.; Klein, T.; Sauer, M.; Kanchanawong, P.; Bourg, N.; Cohen, M. M.; Lelandais, B.; Zimmer, C.

2021-09-09 bioinformatics 10.1101/2021.09.08.459385 medRxiv
Top 0.1%
79.1%
Show abstract

Novel insights and more powerful analytical tools can emerge from the reanalysis of existing data sets, especially via machine learning methods. Despite the widespread use of single molecule localization microscopy (SMLM) for super-resolution bioimaging, the underlying data are often not publicly accessible. We developed ShareLoc (https://shareloc.xyz), an open platform designed to enable sharing, easy visualization and reanalysis of SMLM data. We discuss its features and show how data sharing can improve the performance and robustness of SMLM image reconstruction by deep learning.

2
Suite3D: Volumetric cell detection for two-photon microscopy

Haydaroglu, A.; Dodgson, S.; Krumin, M.; Landau, A.; Baruchin, L. J.; Chang, T.; Guo, J.; Meyer, D.; Reddy, C. B.; Zhong, J.; Ji, N.; Schröder, S.; Harris, K. D.; Vaziri, A.; Carandini, M.

2025-04-01 neuroscience 10.1101/2025.03.26.645628 medRxiv
Top 0.1%
78.9%
Show abstract

In two-photon imaging of neuronal activity it is common to acquire 3-dimensional volumes. However, these volumes are typically processed plane by plane, leading to duplicated cells across planes, reduced signal-to-noise ratio per cell, uncorrected axial movement, and missed cells. To overcome these limitations, we introduce Suite3D, a volumetric cell detection pipeline. Suite3D corrects for 3D brain motion, estimating axial motion and improving estimates of lateral motion. It detects neurons using 3D correlation, which improves the signal-to-background ratio and detectability of cells. Finally, it performs 3D segmentation, detecting cells across imaging planes. We validated Suite3D with data from conventional multi-plane microscopes and advanced volumetric microscopes, at various resolutions and in various brain regions. Suite3D successfully detected cells appearing on multiple imaging planes, improving cell detectability and signal quality, avoiding duplications, and running >20x faster than a prior volumetric pipeline. Suite3D offers a powerful solution for analyzing volumetric two-photon data.

3
A browser-based platform for storage, visualization, and analysis of large-scale 3D images in HPC environments

Walker, L. A.; Lee, W. J.; Duan, B.; Weatherspoon, M.; Li, Y.; Shen, F. Y.; Cheng, M.-C.; Niu, X.; Eggerd, J.; Xie, B.; Cheng, J. H. P.; Xu, H.; Zhong, X.-X.; Jiang, H.; Cui, M.; Yan, Y.; Cai, D.

2024-10-04 bioinformatics 10.1101/2024.10.04.616591 medRxiv
Top 0.1%
76.6%
Show abstract

High-throughput microscopy necessitates 3D image storage, visualization, and analysis in terabyte to petabyte scales. Centralized high-performance computing (HPC) infrastructure provides the resources, but interfacing software is limited. We developed an extensible platform with a scalable image storage format (SISF) and content delivery network (CDN) to enable fast, random access to compressed data. Additionally, we built a cloud-based nTracer2 software to annotate neuron morphology from SISF images in a web browser.

4
NaVis: a virtual microscopy framework for interactive, high-resolution navigation of spatial transcriptomics data

Oshinjo, A.; Wu, J.; Petrov, P.; Izzi, V.

2026-02-19 bioinformatics 10.64898/2026.02.18.706509 medRxiv
Top 0.1%
73.6%
Show abstract

Despite the wide adoption of spatial transcriptomics (ST) into the biomedical community, its practical use remains constrained by a fundamental resolution-coverage trade-off and by reliance on computationally intensive and static workflows. As a result, transcriptome-wide spatial data are typically interpreted as ad-hoc processed outputs rather than explored dynamically as one would do with stained or fluorescence tissue images, limiting ST accessibility and slowing biological insight. Here we introduce NaVis, a web-based virtual microscopy framework that redefines how spatial transcriptomics is experienced. NaVis enables near-real-time, on-demand super-resolution inference from low-resolution whole-transcriptome platforms (10x Genomics Visium V1/V2, Cytassist and VisiumHD), generating high-resolution reconstructions that approach microscopy-level detail while preserving transcriptome-wide coverage. Unlike conventional interpolation approaches that produce fixed images, NaVis computes and refines spatial reconstructions interactively as users navigate tissue sections, transforming resolution from a platform-imposed constraint into a dynamic, user-controlled parameter. Also, NaVis is delivered through a fully point- and-click browser interface requiring no coding expertise, thus removing computational mediation and allowing clinicians, pathologists and experimental researchers to directly interrogate spatial molecular architecture. By coupling high-resolution inference with immediate visual interaction, NaVis shifts spatial transcriptomics from a static computational analysis to an exploratory, microscopy-like modality, broadening both its accessibility, conceptual reach, and potential for biological discoveries.

5
High-speed, multi-Z confocal microscopy for voltage imaging in densely labeled neuronal populations

Weber, T. D.; Moya, M. V.; Mertz, J.; Economo, M. N.

2021-12-13 neuroscience 10.1101/2021.12.10.472140 medRxiv
Top 0.1%
71.9%
Show abstract

Genetically encoded voltage indicators (GEVIs) hold great promise for monitoring neuronal population activity, but GEVI imaging in dense neuronal populations remains difficult due to a lack of contrast and/or speed. To address this challenge, we developed a novel confocal microscope that allows simultaneous multiplane imaging with high-contrast at near-kHz rates. This approach enables high signal-to-noise ratio voltage imaging in densely labeled populations and minimizes optical crosstalk during concurrent optogenetic photostimulation.

6
Petabyte-Scale Multi-Morphometry of Single Neurons for Whole Brains

Jiang, S.; Wang, Y.; Liu, L.; Zhao, S.; Chen, M.; Zhao, X.; Peng, X.; Ding, L.; Ruan, Z.; Dong, H.; Ascoli, G. A.; Hawrylycz, M.; Zeng, H.; Peng, H.

2021-01-10 neuroscience 10.1101/2021.01.09.426010 medRxiv
Top 0.1%
71.0%
Show abstract

Recent advances in neuroscience make the extraction of full neuronal morphology at whole brain dataset available. To produce quality morphometry at large scale, it is highly desirable but extremely challenging to efficiently handle petabyte-scale high-resolution whole brain imaging database. Here, we developed a multi-level method to produce high quality somatic, dendritic, axonal, and potential synaptic morphometry, which was made possible by utilizing necessary petabyte hardware and software platform to optimize both the data and workflow management. Our method also boosts data sharing and remote collaborative validation. We highlight a petabyte application dataset involving 62 whole mouse brains, from which we identified 50,233 somata of individual neurons, profiled the dendrites of 11,322 neurons, reconstructed the full 3-D morphology of more than one thousand neurons including their dendrites and full axons, and detected million scale putative synaptic sites derived from axonal boutons. Analysis and simulation of these data indicate the promise of this approach for modern large-scale morphology applications.

7
LivecellX: A Deep-learning-based, Single-Cell Object-Oriented Framework for Quantitative Analysis in Live-Cell Imaging

Ni, K.; Yu, G.; Zheng, Z.; Lu, Y.; Poe, D.; Zhang, S.; Wang, Z.; Khurana, Y.; Lu, Y.; Chen, Y.; Zhou, S.; Sanborn, M.; Wang, W.; Xing, J.

2025-05-14 biophysics 10.1101/2025.02.23.639532 medRxiv
Top 0.1%
70.8%
Show abstract

Live-cell imaging uniquely captures single-cell dynamics in space and time, but robust analysis is limited by segmentation and tracking errors that accumulate across frames. We present LivecellX, a deep-learning-based pipeline that integrates instance-level segmentation error correction with trajectory refinement, leveraging temporal context to recover accurate cell tracks. LivecellX also introduces a benchmark dataset with detailed annotations of common error classes, providing a resource for method development and evaluation. Beyond error correction, the framework incorporates modules for classifying biological processes, reconstructing cell lineages, and analyzing dynamic behaviors. Users can interact with the system programmatically or through a Napari-based graphical interface, enabling flexible integration into diverse workflows. By coupling error-aware correction with comprehensive lineage and dynamics analysis, LivecellX establishes an open, extensible platform that advances the accuracy and scalability of live-cell imaging studies.

8
CDeep3M-Preview: Online segmentation using the deep neural network model zoo

Haberl, M. G.; Wong, W.; Penticoff, S.; Je, J.; Madany, M.; Borchardt, A.; Boassa, D.; Peltier, S.; Ellisman, M. H.

2020-03-26 neuroscience 10.1101/2020.03.26.010660 medRxiv
Top 0.1%
69.9%
Show abstract

Sharing deep neural networks and testing the performance of trained networks typically involves a major initial commitment towards one algorithm, before knowing how the network will perform on a different dataset. Here we release a free online tool, CDeep3M-Preview, that allows end-users to rapidly test the performance of any of the pre-trained neural network models hosted on the CIL-CDeep3M modelzoo. This feature makes part of a set of complementary strategies we employ to facilitate sharing, increase reproducibility and enable quicker insights into biology. Namely we: (1) provide CDeep3M deep learning image segmentation software through cloud applications (Colab and AWS) and containerized installations (Docker and Singularity) (2) co-hosting trained deep neural networks with the relevant microscopy images and (3) providing a CDeep3M-Preview feature, enabling quick tests of trained networks on user provided test data or any of the publicly hosted large datasets. The CDeep3M-modelzoo and the cellimagelibrary.org are open for contributions of both, trained models as well as image datasets by the community and all services are free of charge.

9
Triplet tumbling microscopy enables in situ quantification of protein complex assembly and dynamics

Lazzari-Dean, J. R.; Millett-Sikking, A.; Rao, P.; Jensvold, Z. D.; Baddock, H.; Ingaramo, M.; Nile, A. H.; York, A. G.; Preciado Lopez, M.

2026-05-11 biophysics 10.64898/2026.05.07.723557 medRxiv
Top 0.1%
66.5%
Show abstract

Protein-protein interactions (PPIs) mediate diverse cellular processes, but PPIs are typically characterized using reconstituted in vitro biochemical and biophysical approaches. Current approaches for PPI detection in living cells are limited in the scope of interactions they can capture and often require prior knowledge of the interacting partners. To close this gap, we developed triplet tumbling microscopy (TTM), which reveals the interactions of a tagged protein of interest in cells in real time. TTM reports protein complex size from rotational diffusion ("tumbling") by leveraging infrared-triggerable emission from triplet states to track tumbling over nanoseconds to hundreds of microseconds. These long-lived triplets overcome the size limitations of existing rotational diffusion-based approaches, enabling TTM to measure species from small protein complexes to organelle-scale beads. In living cells, we apply TTM to detect PPIs, quantify fraction bound, and distinguish protein complexes by size. We measure diverse types of interactions, including rapamycin-induced dimerization, p53 homo-oligomerization, and binding of the E3-ligase E6AP to the human papilloma virus 16 E6 protein. The required hardware is compatible with most fluorescent microscopes, making TTM a versatile way to extract molecular insights from the complex context of living cells. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=109 SRC="FIGDIR/small/723557v1_ufig1.gif" ALT="Figure 1"> View larger version (27K): org.highwire.dtl.DTLVardef@1e70768org.highwire.dtl.DTLVardef@974813org.highwire.dtl.DTLVardef@1fd122borg.highwire.dtl.DTLVardef@1b3da96_HPS_FORMAT_FIGEXP M_FIG C_FIG

10
Neural space-time model for dynamic scene recovery in multi-shot computational imaging systems

Cao, R.; Divekar, N. S.; Nunez, J.; Upadhyayula, S.; Waller, L.

2024-01-20 biophysics 10.1101/2024.01.16.575950 medRxiv
Top 0.1%
65.3%
Show abstract

Computational imaging reconstructions from multiple measurements that are captured sequentially often suffer from motion artifacts if the scene is dynamic. We propose a neural space-time model (NSTM) that jointly estimates the scene and its motion dynamics. Hence, we can both remove motion artifacts and resolve sample dynamics. We demonstrate NSTM in three computational imaging systems: differential phase contrast microscopy, 3D structured illumination microscopy, and rolling-shutter DiffuserCam. We show that NSTM can recover subcellular motion dynamics and thus reduce the misinterpretation of living systems caused by motion artifacts.

11
Convpaint - Universal framework for interactive pixel classification using pretrained neural networks

Hinderling, L.; Witz, G.; Schwob, R.; Stojiljkovic, A.; Dobrzynski, M.; Vladymyrov, M.; Frei, J.; Graedel, B.; Frismantiene, A.; Pertz, O.

2024-09-14 bioinformatics 10.1101/2024.09.12.610926 medRxiv
Top 0.1%
62.5%
Show abstract

We develop Convpaint, a universal computational framework for interactive pixel classification. Convpaint utilizes pretrained convolutional neural networks (CNNs) or vision transformers (ViTs) for feature extraction and enables easy segmentation across a wide variety of tasks. Available within the Python-based napari software ecosystem, Convpaint integrates seamlessly with other plugins into image processing pipelines, which we demonstrate with three workflows across different data modalities.

12
2D, or not 2D? Investigating Vertical Signal Integrity of Tissue Slices

Tiesmeyer, S.; Müller-Bötticher, N.; Malt, A.; Long, B.; Marco Salas, S.; Kiessling, P.; Horn, P.; Guillot, A.; Kümmerle, L.; Ma, L.; Tacke, F.; Theis, F. J.; Kuppe, C.; Nilsson, M.; Eils, R.; Ishaque, N.

2025-01-15 bioinformatics 10.1101/2025.01.13.632601 medRxiv
Top 0.1%
61.8%
Show abstract

Imaging-based spatially resolved transcriptomics can localize transcripts within tissue sections in 3D. Cell segmentation assigns transcripts to cells and precedes annotation of cell function. However, cell segmentation is usually performed in 2D, thus unable to deal with spatial doublets arising from overlapping cells, resulting in segmented cells containing transcripts originating from multiple cell types. Here we present a computational tool called ovrlpy that identifies overlapping cells, tissue folds, and inaccurate cell-segmentation by analyzing transcript localization in 3D.

13
Collaborative Augmented Reconstruction for Scaled Production of 3D Neuron Morphology in Mouse and Human Brains

Zhang, L.; Huang, L.; Yuan, Z.; Hang, Y.; Zeng, Y.; Li, K.; Wang, L.; Zeng, H.; Chen, X.; Zhang, H.; Xi, J.; Chen, D.; Gao, Z.; Le, L.; Chen, J.; Ye, W.; Liu, L.; Wang, Y.; Peng, H.

2023-10-07 neuroscience 10.1101/2023.10.06.561172 medRxiv
Top 0.1%
61.5%
Show abstract

Digital reconstruction of the intricate 3D morphology of individual neurons from microscopic images is widely recognized as a crucial challenge in both individual research laboratories and large-scale scientific projects focusing on cell types and brain anatomy. This task often fails both conventional manual reconstruction and state-of-the-art automatic reconstruction algorithms, even many of which are based on artificial intelligence (AI). It is also critical but challenging to organize multiple neuroanatomists to produce and cross-validate biologically relevant and agreeable reconstructions in scaled data production. Here we propose an approach based on collaborative human intelligence augmented by AI. Specifically, we have developed a Collaborative Augmented Reconstruction (CAR) platform for neuron reconstruction at scale. This platform allows for immersive interaction and efficient collaborative-editing of neuron anatomy using a variety of client devices, such as desktop workstations, virtual reality headsets, and mobile phones, enabling users to contribute anytime and anywhere and take advantage of several AI-based automation tools. We have tested CARs applicability for challenging mouse and human neurons towards a scaled and faithful data production. Our data indicate that the CAR platform is suitable for generating tens of thousands of neuronal reconstructions used in our companion studies.

14
BeadBuddy: user-friendly, nanometer-scale registration of single-molecule imaging data

Lionnet, T.; Clark, F. T.; Whitney, P. H.; Saiz, N.; Ziarno, A.

2025-11-29 biophysics 10.1101/2025.11.25.690467 medRxiv
Top 0.1%
60.5%
Show abstract

Single-molecule localization microscopy (SMLM) captures nanoscale detail with fluorescence imaging. As large-area cameras and multiplex imaging become standard, chromatic errors are growing more complex, challenging precise analysis. To address this challenge, we present BeadBuddy, an open-source, user-friendly software that uses images of fluorescent beads to model and correct 3D, spatially varying chromatic errors. BeadBuddy achieves sub-voxel resolution in DNA Fluorescence In Situ Hybridization (FISH) and is applicable across SMLM modalities.

15
Leonardo: a toolset to correct sample-induced artifacts in light sheet microscopy images

Liu, Y.; Mueller, G. F.; Kowitz, L.; Chobola, T.; Weiss, K.; Maier, P.; Luo, J.; Roessing, M.; Stenzel, M.; Grueneboom, A.; Paetzold, J.; Erturk, A.; Navab, N.; Marr, C.; Chen, J.; Huisken, J.; Peng, T.

2025-10-27 bioinformatics 10.1101/2025.10.26.684661 medRxiv
Top 0.1%
60.1%
Show abstract

Selective plane illumination microscopy (SPIM, also known as light sheet fluorescence microscopy) is the method of choice for studying morphogenesis and function in biological specimens over extended periods, as it permits gentle and rapid volumetric imaging. In inhomogeneous samples, however, sample-induced artifacts, including light absorption, scattering, and refraction, can impact the image quality, particularly as the focal plane gets deeper into the sample. Here, we present Leonardo, the first toolbox designed to address the major sample-induced artifacts by using two modules: (1) DeStripe removes stripe artifacts in SPIM caused by light absorption while preserving fine sample structures; (2) Fuse reconstructs a single high-quality image from dualsided illumination and/or dual-sided detection, while eliminating blur and optical distortions caused by light scattering and refraction. The efficacy of Leonardo is validated on a wide range of biological samples, from minimally invasive experiments on sensitive specimens (translucent embryonic and optically opaque larval zebrafish) to cleared mouse samples up to two centimeters in size. We provide model code and a Napari-based graphical user interface, enabling the SPIM community to easily apply Leonardo to advance light sheet imaging of inhomogeneous and complex specimens.

16
Denoising-based Image Compression for Connectomics

Minnen, D.; Januszewski, M.; Shapson-Coe, A.; Schalek, R. L.; Balle, J.; Lichtman, J. W.; Jain, V.

2021-05-30 neuroscience 10.1101/2021.05.29.445828 medRxiv
Top 0.1%
60.0%
Show abstract

Connectomic reconstruction of neural circuits relies on nanometer resolution microscopy which produces on the order of a petabyte of imagery for each cubic millimeter of brain tissue. The cost of storing such data is a significant barrier to broadening the use of connectomic approaches and scaling to even larger volumes. We present an image compression approach that uses machine learning-based denoising and standard image codecs to compress raw electron microscopy imagery of neuropil up to 17-fold with negligible loss of 3d reconstruction and synaptic detection accuracy.

17
A quantitative pipeline for whole-mount deep imaging and multiscale analysis of gastruloids

Gros, A.; Vanaret, J.; Dunsing-Eichenauer, V.; Rostan, A.; Roudot, P.; Lenne, P.-F.; Guignard, L.; Tlili, S.

2024-08-16 biophysics 10.1101/2024.08.13.607832 medRxiv
Top 0.1%
59.9%
Show abstract

Whole-mount 3D imaging at the cellular scale is a powerful tool for exploring complex processes during morphogenesis. In organoids, it allows examining tissue architecture, cell types, and morphology simultaneously in 3D models. However, cell packing in multilayered organoid tissues hinders both deep imaging and quantification of cell-scale processes. To address these challenges, we developed an experimental and computational pipeline to extract properties at scales ranging from cell to tissue. The experimental module is based on two-photon imaging of immunostained organoids. The computational module corrects for optical artifacts, performs accurate 3D nuclei segmentation and reliably quantifies gene expression. We provide the computational module as a user-friendly Python package called Tapenade, along with napari plugins which enable joint data processing and exploration across scales. We demonstrate the pipeline by quantifying 3D spatial patterns of gene expression and nuclear morphology in gastruloids, revealing how local cell deformations and gene co-expression relate to tissue-scale organization. This quantitative pipeline improves our understanding of gastruloid development, and lays the groundwork for a wide range of multi-layered organoids and tumoroids systems

18
CellBin:a generalist framework to process spatial omics data to cell level

Zhang, Y.; Liu, H.; Wang, H.; Li, Z.; Li, Y.; Chen, J.; Fan, J.; Yi, J.; Shi, C.; Ren, X.; Kang, Q.; Bai, Y.; Fang, S.; Guo, J.; Heng, Y.; Jia, D.; Liao, S.; Chen, A.; Shao, H.; Li, M.

2025-10-28 bioinformatics 10.1101/2025.10.23.683357 medRxiv
Top 0.1%
59.9%
Show abstract

Spatial omics has rapidly expanded with increasingly diverse imaging modalities, molecular targets, and chip sizes. However, no general framework currently exists to construct cell level matrices that are robust across platforms and omics types. Here we present CellBin, a universal and scalable frame-work that unifies image stitching, cell segmentation, and spot-to-cell mapping for multiple spatial omics technologies. CellBin integrates a multi-field weighted stitching algorithm for large-area images, a family of U-Net-based models trained across diverse staining modalities, and an optimized computational architecture for high-throughput processing. Across five technological platforms and three omics data types, CellBin achieves robust segmentation and accurate single-cell matrix construction, consistently outperforming seven state-of-the-art methods in F1-score, cell size precision, and annotation accuracy. By providing a generalizable, cross-platform solution, CellBin bridges multiple spatial omics, enabling unified, high-resolution cell level analyses across technologies.

19
Multimodal alignments of in vivo imaging and spatial biology datasets at cellular resolution

Wang, L.; Jiang, X.; Sun, X.; Chattree, G. M.; Cetin, A.; Cai, X.; Paul, E.; Chrapkiewicz, R.; Hernandez, O.; Ke, Y.; Yoda, T.; Dinc, F.; Kurtkaya, B.; Zhang, Y.; Zhang, Z.; Schnitzer, M. J.

2026-05-01 neuroscience 10.64898/2026.04.28.719500 medRxiv
Top 0.1%
59.0%
Show abstract

Parallel revolutions in intravital microscopy and spatial biology techniques have respectively enabled large-scale recordings of cellular dynamics in live animals and multi-dimensional molecular profiling at single-cell resolution. However, due to the challenges of aligning data from different modalities at cellular resolution, these two transformational approaches have generally been applied on separate biological samples, stymying the ability to link activity patterns and molecular attributes in the same exact cells. To enable routine, multimodal investigations of cells in vivo dynamics and molecular content, we created TRU-FACT (Total Registration Under Functional Activity, Connectivity, and Transcriptomics), a broadly applicable experimental and computational pipeline for registering large populations of individual cells across intravital imaging and spatial biology datasets. The pipeline combines three key innovations: an optomechanical tissue handling and alignment method to parallelize specimen planes, a graph-theoretic method to register individual cells based on their geometric relationships to neighboring cells, and a statistical framework that provides for each cell an a posteriori probability of correct registration. We validated TRU-FACT with several preparations for imaging neural Ca2+ activity in cortical and deep brain areas in head-fixed and freely behaving mice, RNA-barcode-expressing viruses for labeling neural projections, and low- and high-plex spatial transcriptomic methods. In mice performing a skilled reaching task, TRU-FACT alignments revealed the movement-related signaling patterns of intratelencephalic, extratelencephalic, and striatum-, superior colliculus-, and thalamus-projecting motor cortical neurons. Overall, TRU-FACT constitutes a scalable, multimodal discovery platform that is applicable to diverse tissue-types and spatial biology techniques, thereby enabling multiscale analyses of many complex biological systems.

20
Sensitive clustering of protein sequences at tree-of-life scale using DIAMOND DeepClust

Buchfink, B.; Ashkenazy, H.; Reuter, K.; Kennedy, J. A.; Drost, H.-G.

Top 0.1%
58.3%
Show abstract

The biosphere genomics era is transforming life science research, but existing methods struggle to efficiently reduce the vast dimensionality of the protein universe. We present DIAMOND DeepClust, an ultra-fast cascaded clustering method optimized to cluster the 19 billion protein sequences currently defining the protein biosphere. As a result, we detect 1.7 billion clusters of which 32% hold more than one sequence. This means that 544 million clusters represent 94% of all known proteins, illustrating that clustering across the tree of life can significantly accelerate comparative studies in the Earth BioGenome era.